Skip to content

feat: confidence-gap safeguards before wet-lab synthesis - #15

Merged
cschanhniem merged 1 commit into
mainfrom
feat/confidence-gaps
Jun 27, 2026
Merged

cschanhniem merged 1 commit into
mainfrom
feat/confidence-gaps

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

Three safeguards added to close critical confidence gaps before committing to the $10k wet-lab synthesis budget.

Gap 1 — Safety scorer v0.4: hydrophobic moment as hemolysis signal

Pilot-panel SEED-005 variants had μH 0.69–0.87 but were trivially scored safety=1.0 because their hydrophobic fraction sat below the old 0.65 threshold. Amphipathicity (μH) is the primary predictor of non-selective membrane disruption per Dathe & Wieprecht (1999).

  • Added: if mu_h > 0.55: risk += (mu_h - 0.55) * 1.5
  • Effect: SEED-005 variants drop to safety 0.52–0.77; SEED-003 variants (shorter, highly charged, moderate μH ≈ 0.59) stay at 0.94
  • Known blind spot documented: melittin-like peptides (bent-helix hemolytic character) are not captured by 1D μH

Gap 2 — Retrospective AUROC benchmark (critical gate)

Compares 44 known AMPs against 44 composition-identical shuffled decoys. Since all composition features (charge, hydrophobic fraction, Boman, GRAVY, length) are identical per pair, only order-dependent features (hydrophobic_moment) can separate them. This tests whether the scoring signal is real.

Result: AUROC = 0.5305 (gate: >0.70 proceed, 0.55–0.70 caution, <0.55 STOP)

This is a hard finding. The model sits in the POOR zone — it is near-random at discriminating known AMPs from their composition-matched shuffles. Root cause: hydrophobic_moment contributes only ~15% of the activity score; the remaining ~85% is composition-based (identical for each AMP/decoy pair).

Implication: the current scoring model's order-sensitive discriminative power is weak. Run make validate-scoring for the full report.

Gap 3 — External predictor checklist

Generates a pilot FASTA + structured markdown table for manual submission to three independent published tools:

  • CAMPR4 (SVM+RF+ANN+DT ensemble)
  • AMPScanner v2.0 (LSTM)
  • dbAMP 2.0 (Random Forest)

Decision gate: ≥12/20 agree → synthesise; 6–11 → wave-1 only; <6 → STOP. Run make external-predict.

New CLI commands

make validate-scoring     # retrospective AUROC (must run before synthesis)
make external-predict     # FASTA + submission checklist
make pilot-confident KEEP=ID1,ID2,...  # filter to externally confirmed candidates

What to do next (honest)

The AUROC=0.5305 result means the scoring model is not sufficiently discriminative in order-sensitive features. Before spending $10k:

  1. Mandatory: Submit pilot FASTA to all 3 external tools (make external-predict). If ≥12/20 agree, the external consensus overrides the weak internal signal.
  2. Recommended: Improve the model by increasing weight of hydrophobic_moment or adding position-specific features (N-terminal charge patch, motif detection).
  3. Alternative: Use a published QSAR/ML AMP classifier (CAMPR4 offline, iAMPpred) as a primary scorer rather than purely physicochemical heuristics.

Test plan

  • 364 tests pass (all green)
  • 23 new tests: test_safety_moment.py (7 tests), test_retrospective.py (16 tests)
  • Pre-push CI hook ran and passed
  • make validate-scoring runs cleanly and outputs AUROC=0.5305 with interpretation
  • Human: submit outputs/pilot_panel.fasta to CAMPR4, AMPScanner v2, dbAMP 2.0 and record results in checklist

🤖 Generated with Claude Code

Three gaps closed to prevent wasting the $10k synthesis budget:

Gap 1 — Safety scorer v0.4: add hydrophobic moment (μH) as a hemolysis
signal. Pilot-panel SEED-005 variants had μH 0.69–0.87 but were trivially
scored safety=1.0 because their hydrophobic fraction sat below the old
0.65 threshold. Now μH > 0.55 adds a risk penalty proportional to the
excess, reducing those candidates to safety 0.52–0.77.

Gap 2 — Retrospective AUROC benchmark: compare 44 known AMPs against
44 composition-matched shuffled decoys (RNG seed=42). Only order-dependent
feature (hydrophobic_moment) can separate them. Gate: >0.70 proceed,
0.55–0.70 caution, <0.55 STOP. Current result: AUROC=0.5305 (POOR —
near-random). Reported honestly via `make validate-scoring`.

Gap 3 — External predictor checklist: generates pilot FASTA and a
structured markdown table for manual submission to CAMPR4, AMPScanner v2,
and dbAMP 2.0. Decision gate: ≥12/20 agree → synthesise; 6–11 → wave-1
only; <6 → STOP. Run via `make external-predict`.

Tests: 23 new tests in test_safety_moment.py and test_retrospective.py.
Full suite: 364 passed.
@cschanhniem
cschanhniem force-pushed the feat/confidence-gaps branch from 47df7cb to 417cec1 Compare June 27, 2026 15:56
@cschanhniem
cschanhniem merged commit 5a24fa1 into main Jun 27, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant